Papers with corpus creation
Alignment Annotation for Clinic Visit Dialogue to Clinical Note Sentence Language Generation (2020.lrec-1)
Copied to clipboard
| Challenge: | Despite advances in natural language processing, converting a clinic visit conversation into a clinical note is a largely unexplored area of research. |
| Approach: | They propose an annotation methodology that is content- and technique- agnostic while associating note sentences to sets of dialogue sentences. |
| Outcome: | The proposed method is content- and technique-agnostic while associating note sentences to sets of dialogue sentences. |
Phonetically Balanced Code-Mixed Speech Corpus for Hindi-English Automatic Speech Recognition (L18-1)
Copied to clipboard
Ayushi Pandey, Brij Mohan Lal Srivastava, Rohit Kumar, Bhanu Teja Nellore, Kasi Sai Teja, Suryakanth V. Gangashetty
| Challenge: | a phonetic balance in code-mixed Hindi-English corpus has been created . code-switching is a common phenomenon in multilingual and bilingual communities . |
| Approach: | They propose to create a phonetically balanced read speech corpus of code-mixed Hindi-English . they use a method to select sentences that contain triphones lower in frequency than a threshold . |
| Outcome: | The proposed corpus is phonetically balanced with a large code-mixed reference corpus. |
ViNLI: A Vietnamese Corpus for Studies on Open-Domain Natural Language Inference (2022.coling-1)
Copied to clipboard
| Challenge: | a large-scale corpus is needed for studies on natural language inference (NLI) for Vietnamese, which can be considered a low-resource language. |
| Approach: | They propose a corpus for evaluating Vietnamese natural language inference models . they use a human-annotated corpus extracted from more than 800 online news articles . |
| Outcome: | The ViNLI corpus is created and evaluated with a strict process of quality control . the best system performance is still far from human performance (a 14.20% gap in accuracy). |